Back

Computer Methods and Programs in Biomedicine

Elsevier BV

Preprints posted in the last 30 days, ranked by how well they match Computer Methods and Programs in Biomedicine's content profile, based on 28 papers previously published here. The average preprint has a 0.04% match score for this journal, so anything above that is already an above-average fit.

1
The dynamics of arterial pressure itself predict intraoperative hypotension beyond its current value: an interpretable additive model validated in 3,069 external patients under a selection-bias-resistant protocol

Oyarzun, R.; Hernandez, P.

2026-08-31 anesthesia 10.64898/2026.08.26.26361468 medRxiv
Top 0.1%
15.6%
Show abstract

Background. Whether predictors of intraoperative hypotension (IOH) add information beyond the mean arterial pressure (MAP) already displayed on the monitor is contested: selection bias in common evaluation designs inflates apparent performance, and the field has called for comparisons against simple MAP-based references under bias-resistant protocols. Existing predictors also depend on proprietary waveform analysis or pulse-contour monitors, restricting both deployment and external validation. Methods. Using 807 non-cardiac surgery patients from the open VitalDB database, we derived an additive gradient boosting model (one split per tree: a learned shape function per variable, no interactions) from three variables computable from an arterial line alone: current MAP, its drift from the patient's own 20-minute baseline, and the growth of its rolling variance (critical slowing down). Evaluation used patient-level 5-fold cross-validation under a strict protocol - exclusion of the 65-75 mmHg grey zone and of all samples already hypotensive at prediction time - with MAP alone (same learner class) as comparator. The frozen model was then validated, without any refitting, on an independent cohort from another continent (MOVER, University of California Irvine) following a pre-registered plan sealed before external data access. Results. In development the pressure-only model reached AUROC 0.907 vs. 0.884 for MAP alone (Delta AUROC +0.023, 95% CI +0.017 to +0.029) at 5 min, with +0.031 and +0.032 at 10 and 15 min, and good calibration (Brier skill +0.418 vs. prevalence). In external validation on 3,069 patients (442,194 samples, 1-minute charting, event prevalence 5.8%), the advantage not only transferred but was larger than in development: AUROC 0.696 vs. 0.638, Delta AUROC +0.058 (95% CI +0.051 to +0.064), meeting both pre-registered gates. Discrimination transferred; calibration did not (external Brier skill -0.014), requiring local recalibration. In the unrestricted scenario, where samples already at threshold are retained, the advantage collapsed (+0.007), reproducing the selection effect this paper documents. A secondary model adding pulse-contour cardiac output and stroke volume variation improved development discrimination further (Delta AUROC +0.035) but could be externally validated in only 39 patients, because those signals are rarely recorded. Conclusions. The dynamics of arterial pressure itself - drift from a patient-specific baseline and variance growth - carry predictive information beyond its current value, in a fully interpretable additive model that requires only an arterial line, no waveform access and no proprietary hardware. The advantage is confirmed in a pre-registered frozen-model external validation of over three thousand patients, and is largest at coarse recording cadence, where instantaneous pressure is least informative.

2
Artificial Intelligence Model: Optimizing Cancer Risk Level Predictions Using Machine learning and deep learning approaches

Abd Aziz, A. B.; Arabiat, A.; Abu Owida, H.; Abuowaida, S.; Alshdaifa, N.; A. Mashagba, H.

2026-08-25 cancer biology 10.64898/2026.08.20.745910 medRxiv
Top 0.2%
3.4%
Show abstract

This study emphasizes the potential of computational techniques in cancer risk assessment, lighting opportunities for specific and data-driven healthcare solutions. This study examines the use of artificial intelligence (AI), machine learning (ML), and deep learning (DL) approaches to improve cancer risk assessment using a Kaggle dataset. The study uses Java-based ML software to create and evaluate multiple predictive models, taking advantage of its powerful libraries and frameworks for processing and analyzing cancer risk indicators. This work analyzes model performance using 10-fold cross-validation, resulting in reliable generalization and accuracy estimates. Several classification techniques, such as Random Forest (RF) logistic regression (LR), decision trees (DT), Naive Bayes (NB), and Multi-layer perceptron (MLP), are used to assess their efficacy in predicting risk levels for various cancer types. To measure classification effectiveness, key performance metrics such as accuracy, precision, recall, and F1 score are produced, in addition to multi-class confusion matrices. The results show that the RF model is the best classifier for classification, with accuracy of 99.85%, F-measure of 99.80%, precision of 99.80%, and sensitivity of 99.90%. These findings demonstrate the model's ability to effectively estimate cancer risk levels among individuals. of cancer risk estimations, allowing for earlier discovery and more effective medical care.

3
An image-based framework for in silico trials of ablation strategies in scar-related ventricular tachycardia

Biasi, N.; Parollo, M.; Vultaggio, D. M.; Zucchelli, G.; Tognetti, A.

2026-08-06 bioengineering 10.64898/2026.08.06.743176 medRxiv
Top 0.2%
3.2%
Show abstract

Scar-related ventricular tachycardia (VT) is sustained by patient-specific structural and functional remodeling of the ventricular substrate, and the optimal substrate-based ablation strategy remains debated. We present an image-based computational framework for conducting controlled in silico trials of VT ablation strategies. Patient-specific left ventricular electrophysiology models were generated from late gadolinium enhancement cardiac magnetic resonance images by incorporating image-derived scar, border-zone tissue with structural fibrosis, fiber orientation, and physiologically plausible Purkinje-driven sinus activation. A dedicated standalone graphical user interface was developed to perform interactive virtual ablation based on imaging-derived or simulated electrophysiological data. We implemented standardized VT reinducibility testing to compare different lesion sets in terms of residual VT inducibility and ablation burden. As a proof of concept, the framework was applied to 20 patients with ischemic or non-ischemic cardiomyopathy undergoing VT ablation. Sustained VT was inducible in 17 patients, yielding 127 sustained VT episodes and 88 unique reentrant circuits at baseline. Four substrate-based ablation strategies were compared: scar homogenization, primary deceleration-zone ablation, primary plus secondary deceleration-zone ablation, and CMR-guided scar dechanneling. All strategies significantly reduced VT inducibility compared with baseline. Scar homogenization achieved the largest reduction in residual unique sustained VTs but required the largest ablated myocardial volume. Conversely, CMR- guided scar dechanneling showed the most favorable efficiency profile by reducing VT inducibility while limiting ablated viable myocardium. The proposed framework enables quantitative comparison of ablation efficacy, ablation burden, and mechanisms of ablation success or failure in image-guided VT therapy planning.

4
FoxTail: An R-Peak-Anchored Event Domain for Visualizing and Quantifying Changes in ECG Dynamics

Garcia, N. M.

2026-08-18 cardiovascular medicine 10.64898/2026.08.16.26360545 medRxiv
Top 0.2%
3.1%
Show abstract

Conventional electrocardiography is highly effective for waveform and rhythm diagnosis, but it is less suited to showing how the internal shape of hundreds or thousands of consecutive heartbeats changes over time. We introduce FOXTAIL, a complementary view that represents each cardiac cycle as an ordered sequence of changes in signal direction. Overlaying these sequences in a fixed visual field makes beat-to-beat organization visible and allows the density, size, stability, and scale persistence of those changes to be measured. We evaluated the representation in recordings containing normal sinus rhythm, paroxysmal atrial fibrillation, severe heart failure, ventricular tachyarrhythmia, and controlled electrode-motion noise. Paired recordings showed that FOXTAIL descriptors can reveal within-person state changes that are not conveyed by a single average beat. The noise and pre-fibrillation analyses also showed that a dense event pattern is not automatically equivalent to physiological complexity, measurement artifact, or impending disease. FOXTAIL is therefore not proposed as a replacement for the diagnostic ECG or as a new classifier, but as an observation and measurement domain for asking a more basic question: how is the electrical organization of the heart changing from one beat to the next, and which of those changes persist across scale?

5
An interpretable, formally verified point-of-care ultrasound risk equation for difficult videolaryngoscopy: development and internal validation

Oyarzun-Silva, R. A.; Hernandez-Hernandez, P.; Fernandez-Vaquero, M. A.; De Luis-Cabezon, N.

2026-09-02 anesthesia 10.64898/2026.08.28.26361621 medRxiv
Top 0.3%
2.8%
Show abstract

Background. Videolaryngoscopy still requires adjuncts or hyperangulated rescue in a clinically important minority, and bedside screening discriminates modestly. Point-of-care ultrasound (POCUS) of the anterior airway is a promising alternative, but existing prediction models are opaque or assume a pre-specified functional form. We developed and internally validated a parsimonious, fully disclosed POCUS risk equation whose form is recovered from data and whose structural properties are machine-checked by formal proof - to our knowledge the first formally verified clinical risk predictor - following TRIPOD+AI 2024. Methods. In a prospective single-centre, single-operator cohort of 259 adults undergoing elective videolaryngoscopy (no-Easy airway 68/259, 26.3%), Sequentially Thresholded Least Squares with bootstrap stability selection (B=300) screened a 71-term library of nine POCUS features and retained a seven-term logistic equation; a two-term bootstrap-stable model was pre-specified as robustness analysis. Internal validation used 5x10 repeated cross-validation plus temporal and device hold-outs, with pre-specified overfitting and optimism assessments. Five behavioural properties of the deployed equation were machine-checked in Lean 4. Results. Two interactions met the |c|/sigma_c>2 stability criterion: skin-to-epiglottis x skin-to-hyoid-bone distance and tongue volume x sagittal tongue area. The seven-term equation reached a 5x10 cross-validated C-statistic of 0.966 (optimism-corrected 0.968) and held across temporal and device hold-outs (0.94-0.97). Calibration-in-the-large matched prevalence, with cross-validated slope 0.90 attenuating to 0.625 out-of-time; standard recalibration restored 0.92 without loss of discrimination. The pre-specified two-term robustness model reproduced this performance (C-statistic 0.964-0.968; events-per-parameter 34; shrinkage 0.99), confirming the result is not an artefact of the screening stage. Net benefit over a clinical baseline was positive across 10-50% thresholds. All five Lean 4 theorems compiled without sorry. Conclusions. A sparse, formally verified POCUS equation predicts difficult videolaryngoscopy with high internally validated discrimination and quantified, modest overfitting. Because the equation was developed in a single-operator cohort and its inputs are operator-dependent, external validation requires prior harmonisation of the measurement protocol and operator credentialing.

6
Automated Identification of Complex Percutaneous Coronary Intervention from Cardiac Catheterization Reports Using Large Language Models

Bhatt, N.; Warner, F.; Miao, J.; Thakker, R.; Joodi, G.; Cantero-Schaffer, P.; Huang, C.; Krumholz, H.; Murugiah, K.

2026-08-10 cardiovascular medicine 10.64898/2026.08.05.26359802 medRxiv
Top 0.3%
2.5%
Show abstract

Background: Manual abstraction of complex percutaneous coronary intervention (PCI) variables from cardiac catheterization reports is labor-intensive and limits scalable cardiovascular research. Large language models (LLMs) may enable automated extraction of procedural data, but their performance remains uncertain. Methods: We evaluated three open-source LLMs (Llama 3.3 70B, Meditron-7B, and BioMistral-7B) using manually annotated cardiac catheterization reports from three hospitals within Yale New Haven Health system. Models were tasked to identify if a procedure note was a PCI procedure, and extract variables used to classify PCI as complex using predefined criteria, including 3 vessels treated, [≥]3 treated lesions, bifurcation PCI with two stents, chronic total occlusion, [≥]3 stents, and total stent length [≥]60 mm. Results: The evaluation cohort included 1,412 clinical notes of which 596 were PCI procedures. Llama 3.3 70B consistently outperformed both domain-specific models across nearly all extraction tasks. For PCI identification, Llama 3 70B had 100.0% sensitivity, 93.8% specificity, 92.1% positive predictive value, 100.0% negative predictive value, 96.4% accuracy, and an F1 score of 95.9%. For complex PCI classification, among 590 evaluable PCI reports, sensitivity was 97.7%, specificity was 80.1%, positive predictive value was 57.6%, negative predictive value was 99.2%, accuracy was 83.9%, and the F1 score was 72.5%. Variables that were explicitly documented, including stent number, stent length, and adjunctive device use, were extracted with high accuracy, whereas performance was lower for variables requiring contextual reasoning, including lesion counting, bifurcation PCI, and chronic total occlusion. Conclusion: High-capacity open-source LLMs can accurately extract complex PCI variables from free-text catheterization reports, supporting LLM-enabled automated phenotyping to reduce manual abstraction and facilitate scalable cardiovascular research.

7
Computational Evaluation of a Turbulence-like Electrical Activity Hypothesis in Atrial Fibrillation: Substrate Remodeling, Critical Wavelength Transition, and Multi-wavelet Maintenance

Chu, X.; Qiao, Q.; Xu, J.; Wang, X.; Li, M.-M.; Jiang, C.-X.; Tang, R.-B.; Liu, T.; Zhao, X.; Ye, H.; Xu, Z.; Han, K.; Fu, B.; Long, D.-Y.

2026-08-10 cardiovascular medicine 10.64898/2026.08.08.26360016 medRxiv
Top 0.4%
2.1%
Show abstract

BACKGROUND: Atrial fibrillation (AF) remains difficult to explain using a single focal driver or rotor-centered mechanism across disease stages. We tested whether progressive atrial substrate remodeling can drive a critical transition toward turbulence-like, decentralized multi wavelet electrical activity. METHODS: We constructed a controlled two-dimensional atrial reaction-diffusion model with six graded substrate-remodeling stages. We evaluated effective wavelength, theoretical wavelet capacity, AF inducibility, vulnerable-window dynamics, spatial randomness, temporal memory, spectral dispersion, nonlinear indices, virtual ablation response and ERP-prolongation reverse mechanistic testing. RESULTS: Progressive remodeling shortened effective wavelength from 12.0 to 2.4 cm and increased theoretical wavelet capacity from 0.69 to 17.36. Inducibility rose sigmoidally as wavelength shortened, with a model-derived transition near lambda50=4.5 cm. Advanced substrates showed increased wavebreak, spatial randomness, short-memory dynamics, broad spectral dispersion, positive nonlinear indices and resistance to random local ablation. Culprit atrial premature beats within the vulnerable window efficiently triggered AF, whereas counter pacing at 20 to 35 ms reduced inducibility from 52% to 11% in stage 2. CONCLUSIONS: In this controlled model, AF initiation and maintenance were linked to substrate-dependent wavelength, wavelet capacity and vulnerable-window triggering. The model-derived transition provides a testable framework for future high-density mapping, patient30 specific modeling and device-based studies. Key Words atrial fibrillation; turbulence-like electrical activity; substrate remodeling; critical wavelength; multi-wavelet re-entry; vulnerable window; culprit premature atrial beat; counter pacing

8
Novel Large Language Model-Based Detection of Echocardiographic Markers of Right Ventricular Dysfunction

Ekambarapu, L.; Pendyal, A.; Lin, A.; Alwakeel, M.; Rajaratnam, A.

2026-08-31 cardiovascular medicine 10.64898/2026.08.26.26361456 medRxiv
Top 0.4%
1.8%
Show abstract

Background: Unstructured biomedical data, such as echocardiography reports, are rich in information but time consuming to analyze at scale. Rule-based, regular expression-driven terminology mapping can only extract individual variables while large language models (LLMs) offer scalable and clinically meaningful interpretations of heterogeneous disease processes. Right ventricular dysfunction (RVD) is an example of a multifactorial disease state in which key structural and physiologic features are captured both narratively and in structured fields, making it an ideal test case for evaluating whether LLMs can recover complex phenotypes that rules based methods routinely miss. Purpose: To compare an LLM-based extraction method to a conventional rules-based schema for identifying and phenotyping echocardiographic features associated with RVD in a large TTE dataset. Methods: MIMIC-III NOTE2NUM echocardiography reports (n = 45,794) were analyzed using GPT-4o-based LLM extraction deployed within a secure health system enclave and were benchmarked against echocardiographic measurements defined in the MIMIC-III dictionary schema. In MIMIC-III, PH was recorded qualitatively (mild/moderate/severe) based on tricuspid regurgitant (TR) jet velocity and then re-coded as present vs. absent. LLM based extraction defined RVD as (1) RV structural abnormality (>= 1 of hypertrophy, dilation, or wall hypo-/akinesis) or (2) RV pressure/volume overload (>= 2 of the following: estimated right atrial pressure > 8 mmHg, TR jet velocity > 2.8 m/s, fractional area change < 35%, tricuspid annular planar systolic excursion < 17 mm, S' < 9.5 cm/s, or E/e' > 14), with PH defined as estimated pulmonary artery systolic pressure > 35 mmHg or qualitative documentation of PH. Results: LLM extraction identified PH in 15,394 (33.6%), RV pressure/volume overload in 14,449 (31.6%), and RV structural abnormalities in 11,955 (26.1%). Co-occurrence was common: overload + structural changes in 9,380 (20.5%), overload + PH in 9,756 (21.3%), structural changes + PH in 6,183 (13.5%), and all three in 5,620 (12.3%). Using the MIMIC-III dictionary schema, PH prevalence was similar (15,371; 33.6%), but RV overload fields were captured less often (pressure overload 1,357 [3.0%], volume overload 1,128 [2.5%], pressure + volume overload 1,093 [2.4%]; any overload field 3,578 [7.8%]), and RV pressure/volume overload with PH was identified in only 731 (1.6%). Conclusions: LLM-based extraction outperforms rules-based schemas for identifying complex disease states not defined by any single variable. By synthesizing multifactorial signals, LLMs can phenotype RVD with higher fidelity and support population-level assessment. Further validation using multimodality imaging, invasive hemodynamics, and clinical outcome data is needed.

9
Bridging the "Ten Walls" of Japanese Healthcare Data: A Comprehensive Semantic Mapping of JIPAD to HL7 FHIR R4 and Institutional Gap Analysis for the Japanese Health Data Space (JHDS)

Ohno, K.; Hashimoto, S.

2026-08-10 health informatics 10.64898/2026.08.06.26359847 medRxiv
Top 0.5%
1.7%
Show abstract

Background: Japan faces critical challenges in medical data interoperability, conceptualized as the "Ten Walls" obstructing the Japanese Health Data Space (JHDS) [1]. The Japanese Intensive Care Patient Database (JIPAD) - Japan's largest national ICU registry with 151 participating facilities - represents a high-quality critical care dataset that remains isolated from international data ecosystems. Objective: To develop a formal mapping of all 122 JIPAD variables to HL7 FHIR R4, characterize the nature and magnitude of semantic gaps, and assess the feasibility of JIPAD integration into the JHDS. Methods: All 122 JIPAD variables (Data Dictionary v3.7.2; Linkage Items List 20231020) were evaluated using ISO 21564 [8]-based semantic equivalence scoring across three tiers: High (direct FHIR R4 Core mapping), Partial (mapping via JP-Core Implementation Guide extensions [3]), and Low/No Equivalence (structural institutional gap). Semantically identical multi-instance fields (e.g., secondary disease codes x5) were consolidated into single mapping entries, yielding 114 mapping entries. Pseudonymization architecture was characterized from primary documentation. Results: Of 114 mapping entries representing the 122 JIPAD variables, 97 (85.1%) achieved High Equivalence via LOINC/SNOMED CT, and 12 (10.5%) achieved Partial Equivalence via JP-Core extensions, value-set translation, or FHIR R4 Core extension mechanisms - yielding a combined technical feasibility of 95.6% (109/114). Only 5 entries (4.4%) were classified as Low/No Equivalence, all attributable to Japan's proprietary disease classification system (288 adult codes; 165 pediatric codes) embedded in the DPC reimbursement framework, plus one Japan-specific procedure (PMX endotoxin adsorption) absent from international terminology systems. Variable-level mapping details are provided in Supplementary Table S1. Critically, JIPAD employs pseudonymization with record-linkage capability, enabling 99% DPC data matching - demonstrating that technical and design-level barriers to FHIR integration have already been resolved. Conclusion: JIPAD is technically and architecturally ready for FHIR integration at a 95.6% level. The remaining 4.4% barrier is exclusively institutional - rooted in MHLW policy frameworks governing the DPC disease classification system [6] - rather than technical. FHIR integration would further unlock pharmacoepidemiological and social epidemiological research currently inaccessible due to data isolation. As the sole national ICU registry providing high-acuity anchor data unavailable in general health records, JIPAD integration is essential for a clinically meaningful JHDS by 2027.

10
Local retraining mitigates domain shift in sepsis prediction: Lessons from translating a neonatal model to mixed intensive care data

Champeaux, S. A.; Booth, J.; Brown, A.; Sebire, N. J.; Drobnjak, I.; Bowyer, S.

2026-08-21 health informatics 10.64898/2026.08.18.26360666 medRxiv
Top 0.6%
1.5%
Show abstract

Background: Machine learning models leveraging electronic health records (EHRs) can support earlier detection of sepsis in intensive care units (ICUs). However, their clinical utility depends on reproducibility across institutions and patient populations. Building on a published pipeline from the Children's Hospital of Philadelphia (CHOP), this study examines how a neonatal sepsis prediction framework performs and can be adapted to a range of intensive care environments, paediatric, cardiac, and neonatal, at Great Ormond Street Hospital (GOSH). Methods: We extracted de-identified ICU EHR data from GOSH and applied feature derivation, unit harmonisation, and temporal sampling to align with the CHOP dataset used by Masino et al. (2019). Seven classifiers were first evaluated using CHOP-trained weights to characterise cross-domain behaviour and then retrained on local data to assess recoverability and site-specific adaptation. Model discrimination was summarised by AUC and F1, and learning curves were used to explore sample efficiency and bias-variance dynamics. Results: Models achieved strong discrimination on the CHOP neonatal cohort but demonstrated reduced performance when transferred to the mixed GOSH ICU population, reflecting anticipated domain and population shift. Retraining on GOSH data restored discrimination (AUC range 0.69-0.86), with Gradient Boosting (AUC 0.86 vs AUC 0.87 at CHOP) and KNN (AUC 0.80 vs AUC 0.79 at CHOP) models performing comparably to their CHOP benchmarks. DeLong's test confirmed statistically significant gains across all classifiers (p < 0.001). Conclusion: ICU cohort and baseline demographic differences between CHOP and GOSH introduced domain shift that limited direct model transfer. Elements of the original preprocessing pipeline could not be reproduced, further constraining transportability. Yet, retraining on local data restored high discrimination, showing that the modelling framework remains robust when re-estimated in new settings. These results highlight local adaptation as a practical route to recover performance and support safe, generalisable deployment of clinical prediction models in mixed clinical environments.

11
A Guided AI Framework for Customizable and Efficient Harmonisation to the OMOP Common Data Model

Nehra, N.; Swami, R.; Dadi, D.; Mishra, R.; Sharma, U.; Verma, P.; Sen, M.; Dhruw, N. K.; Jha, A. K.

2026-08-12 bioinformatics 10.64898/2026.08.07.742453 medRxiv
Top 0.6%
1.5%
Show abstract

AO_SCPLOWBSTRACTC_SCPLOWGetting clinical data from different sources to "talk" to each other within the OMOP Common Data Model (CDM) is arguably the most tedious part of multi-center research. While this integration is essential, the transformation process is frequently a manual grind, requiring a rare overlap of deep clinical knowledge and technical expertise. In this paper, we present a framework designed to alleviate some of the burden on the researcher by automating data harmonization through two distinct steps: structural schema mapping and terminological standardization. For the structural piece, we moved away from "black box" logic in favor of a stateful workflow managed by large language models (LLMs) and directed acyclic graphs. By profiling EHR data at the source, our system generates context-aware dictionaries that offer ranked mapping suggestions alongside confidence scores. While our benchmarking showed a 97.5% agreement rate at the schema level and an 84% agreement rate at the value level when compared with human experts, the system appears most effective when treated as a "co-pilot" rather than a total replacement for human oversight. To handle value-level standardization, we implemented a hybrid search strategy that pairs the semantic depth of SapBERT embeddings with the literal precision of fuzzy string matching. By using FAISS for rapid similarity retrieval, the engine attempts to resolve messy or "noisy" clinical descriptions to standard OMOP concepts. This approach seems particularly promising for handling the non-standardized labels that often plague smaller, local datasets. Ultimately, our results suggest that this guided approach can shift the timeline for OHDSI-compliant warehousing from weeks of manual curation to a more manageable and scalable pipeline, potentially lowering the barrier to entry for smaller research teams.

12
When can predictive uncertainty be trusted? A methodological evaluation in free-living wearable electrocardiogram signal-quality assessment

Tran, K. D.

2026-08-28 health informatics 10.64898/2026.08.25.26361304 medRxiv
Top 0.6%
1.4%
Show abstract

Uncertainty quantification is proposed as a safeguard for machine-learning systems in health-related signal analysis, but an uncertainty score is useful only if it behaves as a reliability signal. Free-living wearable electrocardiogram (ECG) signal-quality assessment provides a test bed because ambiguity, artifact, and acquisition shift can alter the relationship between confidence and correctness. This study evaluates predictive uncertainty under ambiguity, controlled corruption, and external distribution shift. 32,224 non-overlapping 10-s windows of synchronised single-lead ECG and three-axis accelerometry from 15 subjects in the Brno University of Technology ECG Quality Database were analysed. Two model families were compared: multinomial logistic regression and Classification and Regression Tree (CART), each progressing from a point estimate to a fixed-structure posterior and then a structure posterior. Expected conditional entropy and mutual information were evaluated as designated aleatoric and epistemic uncertainty measures, with max-softmax uncertainty as a confidence baseline. Validation covered error ranking, selective prediction, behavioural probes, posterior structural diversity, recorded-noise stress testing, and zero-shot external transfer. The logistic structure posterior retained an expected 8.5 of nine features and concentrated on near-complete masks, yielding little additional predictive diversity. Bayesian CART produced 221 distinct complete topologies among 238 retained draws and stronger score-dependent selective-risk behaviour. Conditional entropy increased with local class overlap, whereas mutual information increased when training information was reduced, although both showed cross-sensitivity. Under recorded noise, predicted quality severity changed more consistently than uncertainty, while external transfer preserved ordinal severity more reliably than uncertainty ordering. These findings show that posterior richness alone does not establish reliable uncertainty. Model-derived uncertainty should therefore be validated against prespecified ambiguity, information, and shift probes before supporting abstention, reacquisition, or downstream decisions.

13
An Automated Patient Identity Verification Framework for Multimodal Medical Imaging Using Deep Metric Learning and Domain Adaptation

Ueda, Y.; Ishida, T.

2026-08-12 health informatics 10.64898/2026.08.11.26360177 medRxiv
Top 0.8%
1.1%
Show abstract

Purpose: Patient identity management is fundamental to healthcare information systems, as identification inconsistencies can compromise patient safety, data integrity, and clinical workflow efficiency. Reliable linkage of medical images acquired across different imaging modalities remains challenging because of variations in image appearance, acquisition geometry, and imaging characteristics. In this study, we developed an automated patient identity verification framework for multimodal medical imaging using deep metric learning and Data-Augmented Domain Adaptation (DADA). Methods: The proposed framework learned modality-invariant patient representations from labeled source-domain data while leveraging unlabeled target-domain data to mitigate cross-modality distribution shifts. Chest radiographs and computed tomography (CT) scout images obtained under routine clinical conditions were retrospectively collected and used for evaluation. Verification performance was assessed using receiver operating characteristic (ROC) analysis, with the area under the ROC curve (AUC) used as the primary performance metric. Results: The proposed framework achieved consistently high verification performance across all evaluation conditions, with AUC values ranging from 0.9997 to 0.9998. Similarity-score distributions demonstrated distinct separation between same-patient and different-patient image pairs despite substantial differences between imaging modalities. Conclusion: These findings indicate that patient-specific anatomical representations can be preserved across heterogeneous imaging domains through metric learning and domain adaptation. The proposed framework may serve as a practical infrastructure component for patient identity management, multimodal data integration, quality assurance, and patient safety applications within healthcare information systems.

14
Conditional Spatial Classification of Expert-Confirmed Interictal Epileptiform Discharge Epochs: An EEG-ECG Ablation and SHAP Analysis

Plabon, A. M.; Mukit, A.; Neyamul, M.; Jehady, O. F.; Zuba, F. T.; Mina, M. F.; Islam, T.

2026-08-19 bioengineering 10.64898/2026.08.13.744348 medRxiv
Top 0.8%
1.0%
Show abstract

Interictal epileptiform discharges (IEDs) are diagnostically important EEG abnormalities observed between seizures. This study addresses a conditional spatial-classification task where every analyzed four-second epoch had already been reviewed and confirmed by experts as containing an IED, and the model assigned that epoch to one of five predefined scalp-distribution categories (generalized, frontal, temporal, occipital, or centro-parietal). The analysis therefore does not evaluate IED-versus-non-IED detection. After preprocessing, 2,514 IED-labelled epochs were analyzed using identical stratified epoch-level partitions, SMOTE based training, 26 handcrafted features per included channel, and multiple machine-learning classifiers. A staged channel ablation compared 19-channel scalp EEG, 21-channel EEG with ECG, and the complete 29-channel input containing scalp EEG, referential, ECG, and EMG channels. The best EEG-only result was obtained with linear discriminant analysis (88.89% test accuracy). CatBoost achieved 93.25% on EEG with ECG channel and 94.44% with the whole channel set. All eight directly comparable classifiers showed numerically higher test accuracy after ECG channel was added; for CatBoost, the increase was 6.35 percentage points. In the EEG with ECG channel, CatBoost model on ECG channel on right and left arm received respectively 15.79% and 15.12% of normalized global SHAP attribution, and beta-band power was the leading of all features (18.76%). These SHAP values indicate model-specific predictive contributions and do not establish physiological biomarkers, causal autonomic mechanisms, or clinical localization. The findings support a limited methodological conclusion which is ECG-derived features were associated with improved internal epoch-level categorization of expert-confirmed IED epochs. They do not establish IED detection, artifact rejection, independent EMG effects, or generalization to unseen patients.

15
Large Language Models Generate Stigmatizing Language During Reasoning Over Real-World Clinical Data

Yang, Y.; Gu, B.; Hathaway, D. B.; Wyss, R.; Marengo, L.; Gibbons, J. B.; Lyndon, S.; Wu, J.; Chen, Q.; Liu, N.; Wang, P. S.; Celi, L. A.; Bates, D. W.; Lin, J.; Zhou, L.; Yang, J.

2026-08-14 health informatics 10.64898/2026.08.12.26360210 medRxiv
Top 0.9%
0.9%
Show abstract

Stigmatizing language in clinical documentation, which conveys negative stereotypes, attitudes, or judgments toward patients, is a recognized source of documentation bias and is associated with poorer care and adverse health outcomes. Although prior stigma-related research has focused on clinician-written EHR notes, the increasing use of large language model (LLM)-generated documentation in clinical workflows raises new concerns about its potential to reproduce or amplify bias and affect patient safety. In this study, we conducted a large-scale assessment of stigmatizing language in LLM-generated reasoning text on 35 real-world clinical tasks across 107 LLMs. We applied a psychiatrist-validated, natural language processing (NLP) system to detect stigma terms in LLM reasoning text and quantified stigma rates of LLM-generated reasoning texts across 3,745 model-task pairs. Results showed that stigma rates ranged from 0% to 33.33%, with 84.06% of pairs containing stigma terms. Open-source models and reasoning models showed statistically higher stigma rates than proprietary (1.97% vs. 1.60%; p < 0.01) and non-reasoning models (2.35% vs. 1.70%; p < 0.0001), while the stigma rate difference between the general and medical models is not statistically significant (2.00% vs. 1.80%; p = 0.26). Stigma rates of LLM outputs correlated negatively with task accuracy (r = -0.304; p < 0.001) and positively with input clinical-text stigma (r = 0.569; p < 0.001), with 19.76% of model-task pairs amplifying stigma in the original input notes. Applying prompt engineering as a destigmatizing approach helped reduce model stigma rates by as much as 91.91% without affecting the model performance. This study shows that stigmatizing language generation is common but reducible during LLMs' reasoning traces, suggesting that well-implemented approaches for LLM monitoring and destigmatizing will be essential for healthcare systems to implement.

16
A mathematical investigation of the interplay between vasculature and intratumoral cellular heterogeneity during tumor progression

Ghosh, S.; Sadhu, G.; Dalal, D.

2026-08-27 systems biology 10.64898/2026.08.26.747242 medRxiv
Top 1.0%
0.9%
Show abstract

Tumors consist of heterogeneous phenotypic cells, such as normoxic cells, which are highly proliferative, and hypoxic cells, which are less proliferative. Their phenotypic switching depends on tumor microenvironmental factors, such as oxygen and nutrient concentrations supplied by local blood vessels. However, during ongoing angiogenesis, the process of sprouting new blood vessels at the tumor site from pre-existing blood vessels, and how this phenotypic switching affects and impacts tumor growth, remains poorly understood. In this article, we formulate a mathematical model to elucidate the crosstalk between vasculature and tumor cellular heterogeneity during tumor progression. The model results show a strong agreement with the experimental data. Our simulation results demonstrate that ongoing angiogenesis increases tumor growth rate. In addition, we observe that the influence of hypoxic cells on phenotypic switching from normoxic to hypoxic is more pronounced than their influence on the transition from hypoxic to normoxic. Furthermore, we perform a global sensitivity analysis using the Sobol's method to assess the importance of the model's parameters. It highlights that the volume at which blood vessels attain half-maximal rate has the maximum effect on the model.

17
Protocol for the Development and Prospective Evaluation of ASHA Assist India: An AI-Assisted Mobile Platform for Community-Based Stroke Prevention in Rural India

Nayak, K. S.; Nirgude, A. S.; Das, R.

2026-08-11 cardiovascular medicine 10.64898/2026.08.10.26360065 medRxiv
Top 1%
0.8%
Show abstract

Background Stroke remains one of the leading causes of mortality and long-term disability worldwide, with low- and middle-income countries bearing a disproportionate share of the global disease burden. In India, delays in risk identification, fragmented referral pathways, and limited continuity of preventive care present significant challenges, particularly in rural communities. As a frontline health worker Accredited Social Health Activists (ASHAs) are strategically positioned to support community-based stroke prevention; however, existing workflows are frequently constrained by multi-tasking, predominantly paper-based documentation and fragmented digital systems. Advances in mobile health, artificial intelligence along with digital health ecosystem provided by Ayushman Bharat Digital Mission (ABDM) provide an opportunity to strengthen community healthcare through integrated digital platforms. Objective This protocol describes the design, system architecture, and prospective evaluation framework of ASHA Assist India, an integrated AI-assisted mobile health platform intended to support community-based stroke prevention by connecting citizens, ASHA workers, Primary Health Centres (PHCs), and higher levels of healthcare facilities within a unified digital ecosystem. Methods ASHA Assist India has been designed as a modular, cloud-based digital health platform supporting standardized data collection, longitudinal health monitoring, referral management, and AI-assisted clinical decision support. The proposed system comprises four user-facing applications corresponding to citizens, ASHA workers, PHCs, and referral hospitals, integrated through a centralized backend providing authentication, secure data management, interoperability, analytics, and notification services. The AI framework includes three planned analytical modules: (i) population-level stroke risk stratification, (ii) longitudinal stroke risk prediction, and (iii) acute stroke symptom recognition. A prospective implementation study is planned to evaluate platform usability, feasibility, workflow integration, implementation outcomes, and operational performance within routine community healthcare settings. Future validation of the AI modules will be conducted using prospectively collected longitudinal datasets. Expected Impact The proposed platform aims to strengthen community-based stroke prevention by improving digital workflow integration, facilitating coordinated referral pathways, and supporting longitudinal monitoring through the existing healthcare providers at health and wellness centres like ASHA, Community Health Officers (CHOs), ANM, etc. Beyond stroke prevention, the modular architecture is intended to provide a scalable framework for future digital health programmes addressing multiple non-communicable diseases within primary healthcare systems. Publication of this protocol establishes a transparent implementation and evaluation framework that may guide future research, digital health innovation, and implementation science in resource-constrained settings.

18
A Human-in-the-Loop Large Language Model System Based on the Model Context Protocol for Differential Diagnosis from Electronic Medical Records and Literature

Lim, H.; Yi, H.; Yoon, J. Y.; Kwon, H.; Lee, D.; Kim, N.

2026-08-21 health informatics 10.64898/2026.08.18.26359085 medRxiv
Top 1%
0.8%
Show abstract

Diagnostic errors, including misdiagnoses and delayed clinical diagnoses, could affect outcomes of a significant patient population, particularly individuals presenting with rare diseases or non-specific symptoms. From rule-based diagnostic decision supporting systems (DDSS) to large language model (LLM) based tools for clinical reasoning have been developed to address these limitations. However, existing DDSS are often proprietary and difficult to integrate, and recent LLM-based tools remain hindered by operational challenges such as cost, resources constraint, and privacy concerns. Moreover, existing systems interpret electronic medical records (EMR) and generate diagnoses separately, limiting continuous evidence-based analysis and imposing repeated clinician involvement. In this paper, we present DDx-Finder, an open-source framework that leverages Model Context Protocol (MCP) servers for direct EMR and literature access, enabling prompt-driven clinical state extraction and reliable case-report re- trieval via generating searching query by LLM, while addressing limitations related to resource demands and privacy concerns. A clinical case study demonstrates the systems feasibility and its potential to provide accessible, transparent, and systematic differential diagnostic support for complex cases.

19
REFINE: Closing the Loop Between Large Language Models and Symbolic Rules in Clinical NLP

Wang, N.; Kakadiaris, A.; Li, C.; Wang, R.; Ahn, J.; Wang, Y.; Fu, S.

2026-08-17 health informatics 10.64898/2026.08.11.26360118 medRxiv
Top 1%
0.8%
Show abstract

Symbolic clinical natural language processing (NLP) systems remain widely used for extracting clinical concepts from electronic health record (EHR) narratives, but maintaining rule resources requires extensive manual error analysis and rule refinement. This study investigates whether large language models (LLMs) can assist in identifying extraction errors and generating candidate rules to improve symbolic clinical NLP systems. Using error reports derived from a multi-site evaluation of a previously validated symbolic model for cognitive and neuropsychiatric-related clinical concepts, we developed a human-in-the-loop framework, REFINE. The framework first uses LLMs to classify extraction errors and generate explanatory reasoning, which can then be incorporated into prompts for rule generation. Three LLMs (GPT-5.2, GPT-4o, GPT-4o-mini) were evaluated under four prompting conditions. LLM-generated rule sets improved performance compared with the baseline NLP-CAM system, increasing F1-score from 0.37 to 0.58. These findings suggest that LLMs can support scalable rule refinement for symbolic clinical NLP systems.

20
Toward Transportable Acute Kidney Injury Prediction: An Explainable XGBoost Model with Temporal Validation Using MIMIC-IV

Okundaye, D. O.; Isiekwene, C. C.

2026-09-03 health informatics 10.64898/2026.09.01.26360393 medRxiv
Top 1%
0.8%
Show abstract

Acute kidney injury (AKI) is a frequent complication within intensive care units, with its sudden onset often missed. This is especially important because a timely window for intervention is required as delayed detection leads to progressively worse outcomes. Existing machine learning and deep learning models have contributed to closing this gap, but their complexity, requiring hundreds to thousands of features, and lack of generalisation pose a limitation that prevents them from being integrated into clinical workflows across different electronic health-record ecosystems. This study presents a 37-feature XGBoost model trained on the MIMIC-IV dataset with 5.4% positive cases, with hyperparameters optimised via Optuna and probabilities calibrated using isotonic regression, designed for transportability across clinical settings. Validation was conducted internally using a temporal patient-level split simulating prospective deployment, training on 2008-2016 data and testing on 2017-2022 data"External validation was performed on the eICU Collaborative Research Database, a multi-centre dataset spanning 208 US hospitals, using the trained model without retraining. SHAP TreeExplainer was used to provide feature-level explainability for individual predictions. Internal testing yielded an AUROC score of 0.794 for predicting AKI onset within a 12-24 hour window. External validation produced a 0.750 AUROC without retraining. Equitable discrimination was observed across gender, age, chronic kidney disease presence, race, and AKI stages on both datasets, with a 95% internal CI of 0.789-0.799 confirming the model's estimate stability. These results suggest that clinically useful prediction systems are achievable with substantially fewer features than current models require.